Papers with inter-annotator agreement
UrBLiMP: A Benchmark for Evaluating the Linguistic Competence of Large Language Models in Urdu (2026.findings-acl)
Copied to clipboard
| Challenge: | Evaluating how large language models capture grammatical structure of low-resource languages remains underexplored. |
| Approach: | They evaluate a set of 5,696 minimal pairs that contrast grammatical acceptability across ten core syntactic and morpho-syntactical phenomena in Urdu. |
| Outcome: | The proposed framework compares multilingual models with the proprietary model . the proposed framework achieves the highest average accuracy on regular phenomena . |
RiQuA: A Corpus of Rich Quotation Annotation for English Literary Text (2020.lrec-1)
Copied to clipboard
| Challenge: | In literature, spoken interactions between characters are of central importance to the narrative. |
| Approach: | They propose to annotate quotations, including their interpersonal structure, for English literary text. |
| Outcome: | The proposed dataset provides a rich view of dialogue structures not available from other available corpora. |
TextAnnotator: A UIMA Based Tool for the Simultaneous and Collaborative Annotation of Texts (2020.lrec-1)
Copied to clipboard
| Challenge: | Existing annotation tools are not efficient for the annotation of corpora and are not error-free. |
| Approach: | They propose to extend existing annotation tools by evaluating their flexibility and efficiency. |
| Outcome: | The proposed system performs platform-independent multimodal annotations and annotates complex textual structures. |
PDFAnno: a Web-based Linguistic Annotation Tool for PDF Documents (L18-1)
Copied to clipboard
| Challenge: | Currently, linguistic annotation tools for PDF documents focus on plain-text documents. |
| Approach: | They propose a web-based linguistic annotation tool for PDF documents . it offers functions for various types of linguistic annotations directly on PDF . |
| Outcome: | The proposed tool can annotate on PDF documents with named entity, dependency relation, and coreference chain. |
LMUNIT: Fine-grained Evaluation with Natural Language Unit Tests (2025.findings-emnlp)
Copied to clipboard
Jon Saad-Falcon, Rajan Pathe Vivek, William Berrios, Nandita Shankar Naik, Matija Franklin, Bertie Vidgen, Amanpreet Singh, Douwe Kiela, Shikib Mehri
| Challenge: | Using natural language unit tests, language models are costly and noisy, and automated metrics provide only coarse, difficult-to-interpret signals. |
| Approach: | They propose a paradigm that decomposes response quality into explicit, testable criteria and a unified scoring model, LMUnit, which combines multi-objective training across preferences, direct ratings, and natural language rationales. |
| Outcome: | The proposed paradigm significantly improves inter-annotator agreement and enables more effective LLM development workflows. |
HighRES: Highlight-based Reference-less Evaluation of Summarization (P19-1)
Copied to clipboard
| Challenge: | Existing methods for summarizing documents are inconsistent due to the difficulty of manual evaluation. |
| Approach: | They propose a method where summaries are evaluated by multiple annotators against the source document via manually highlighted salient content. |
| Outcome: | The proposed method improves inter-annotator agreement while highlighting differences among systems. |
Human vs. Machine Perceptions on Immigration Stereotypes (2024.lrec-main)
Copied to clipboard
| Challenge: | a growing number of natural language processing models leave aside the language itself . a recent paradigm in the computational linguistics community is training models on specific perspectives of a segment of the population or an individual. |
| Approach: | They propose to use BERT-based classification models to detect stereotypes related to immigrants . they compare models with predictions from GPT-4 and annotated tweets from Spanish Twitter . |
| Outcome: | The proposed models are compared with predictions from the dataset of Spanish Twitter posts containing stereotypes . the models are confident in their predictions and more accurate for implicit stereotypes, the authors show . |
Who Watches the Watchmen? Humans Disagree With Translation Metrics on Unseen Domains (2026.findings-acl)
Copied to clipboard
| Challenge: | Existing studies that analyze unseen domains vary translation systems, annotators, or evaluation conditions, confounding domain effects with human annotation noise. |
| Approach: | They propose to use human error span annotations to evaluate translations of six translation systems across one seen news domain and two unseen technical domains to address these biases. |
| Outcome: | The proposed model improves on the human annotations in two unseen domains and on the news domains. |